Papers by Hawau Olamide Toyin

5 papers
Exploring the Limitations of Detecting Machine-Generated Text (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in the quality of the generation of text by large language models have spurred research into identifying machine-generated text.
Approach: They audit classification performance for detecting machine-generated text by evaluating on texts with varying writing styles.
Outcome: The proposed methods are highly sensitive to stylistic changes and complexity, and in some cases degrade entirely to random classifiers.
Voice of a Continent: Mapping Africa’s Speech Technology Frontier (2025.emnlp-main)

Copied to clipboard

Challenge: linguistic diversity in Africa is underrepresented in speech technologies, creating barriers to digital inclusion.
Approach: They propose a benchmarking framework to map the continent's linguistic diversity and map its impact on downstream African speech tasks.
Outcome: The proposed model achieves state-of-the-art across multiple African languages and speech tasks.
Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages (2026.acl-long)

Copied to clipboard

Challenge: idioms are a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpretation.
Approach: They propose a multilingual idiom dataset that provides idiomatic expressions in both sentence-level and conversational contexts.
Outcome: The proposed model performs well with low-resource idioms, but lacks contextual inference.
Where Are We? Evaluating LLM Performance on African Languages (2025.acl-long)

Copied to clipboard

Challenge: African languages are underrepresented in NLP due to policies that favor foreign languages and create data inequities.
Approach: They integrate theoretical insights on Africa’s language landscape with an empirical evaluation using Sahara datasets.
Outcome: The proposed model improves on a benchmark curated from large-scale, publicly accessible datasets capturing the continent's linguistic diversity.
Dialectal Coverage And Generalization in Arabic Speech Recognition (2025.acl-long)

Copied to clipboard

Challenge: Existing ASR systems cover the modern standard Arabic variety but fail to cover the multitude of spoken variants.
Approach: They propose a suite of automatic speech recognition models optimized to recognize multiple variants of spoken Arabic.
Outcome: The proposed models show coverage and performance gains compared to prior models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations